NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Annotation-free prediction of microbial dioxygen utilization

https://doi.org/10.1128/msystems.00763-24

Flamholz, Avi I; Goldford, Joshua E; Richter, Philippa A; Larsson, Elin M; Jinich, Adrian; Fischer, Woodward W; Newman, Dianne K (October 2024, mSystems)
Greening, Chris (Ed.)
ABSTRACT Aerobes require dioxygen (O₂) to grow; anaerobes do not. However, nearly all microbes—aerobes, anaerobes, and facultative organisms alike—express enzymes whose substrates include O₂, if only for detoxification. This presents a challenge when trying to assess which organisms are aerobic from genomic data alone. This challenge can be overcome by noting that O₂utilization has wide-ranging effects on microbes: aerobes typically have larger genomes encoding distinctive O₂-utilizing enzymes, for example. These effects permit high-quality prediction of O₂utilization from annotated genome sequences, with several models displaying ≈80% accuracy on a ternary classification task for which blind guessing is only 33% accurate. Since genome annotation is compute-intensive and relies on many assumptions, we asked if annotation-free methods also perform well. We discovered that simple and efficient models based entirely on genomic sequence content—e.g., triplets of amino acids—perform as well as intensive annotation-based classifiers, enabling rapid processing of genomes. We further show that amino acid trimers are useful because they encode information about protein composition and phylogeny. To showcase the utility of rapid prediction, we estimated the prevalence of aerobes and anaerobes in diverse natural environments cataloged in the Earth Microbiome Project. Focusing on a well-studied O₂gradient in the Black Sea, we found quantitative correspondence between local chemistry (O₂:sulfide concentration ratio) and the composition of microbial communities. We, therefore, suggest that statistical methods like ours might be used to estimate, or “sense,” pivotal features of the chemical environment using DNA sequencing data.IMPORTANCEWe now have access to sequence data from a wide variety of natural environments. These data document a bewildering diversity of microbes, many known only from their genomes. Physiology—an organism’s capacity to engage metabolically with its environment—may provide a more useful lens than taxonomy for understanding microbial communities. As an example of this broader principle, we developed algorithms that accurately predict microbial dioxygen utilization directly from genome sequences without annotating genes, e.g., by considering only the amino acids in protein sequences. Annotation-free algorithms enable rapid characterization of natural samples, highlighting quantitative correspondence between sequences and local O₂levels in a data set from the Black Sea. This example suggests that DNA sequencing might be repurposed as a multi-pronged chemical sensor, estimating concentrations of O₂and other key facets of complex natural settings.
more » « less
Full Text Available
Protein cost minimization promotes the emergence of coenzyme redundancy

https://doi.org/10.1073/pnas.2110787119

Goldford, Joshua E; George, Ashish B; Flamholz, Avi I; Segrè, Daniel (April 2022, Proceedings of the National Academy of Sciences)

Significance Metabolism relies on a small class of molecules (coenzymes) that serve as universal donors and acceptors of key chemical groups and electrons. Although metabolic networks crucially depend on structurally redundant coenzymes [e.g., NAD(H) and NADP(H)] associated with different enzymes, the criteria that led to the emergence of this redundancy remain poorly understood. Our combination of modeling and structural and sequence analysis indicates that coenzyme redundancy may not be essential for metabolism but could rather constitute an evolved strategy promoting efficient usage of enzymes when biochemical reactions are near equilibrium. Our work suggests that early metabolism may have operated with fewer coenzymes and that adaptation for metabolic efficiency may have driven the rise of coenzyme diversity in living systems.
more » « less
Full Text Available
Environmental boundary conditions for the origin of life converge to an organo-sulfur metabolism

https://doi.org/10.1038/s41559-019-1018-8

Goldford, Joshua E.; Hartman, Hyman; Marsland, Robert; Segrè, Daniel (December 2019, Nature Ecology & Evolution)

Full Text Available
Modern views of ancient metabolic networks

https://doi.org/10.1016/j.coisb.2018.01.004

Goldford, Joshua E.; Segrè, Daniel (April 2018, Current Opinion in Systems Biology)

Full Text Available
Emergent simplicity in microbial community assembly

https://doi.org/10.1126/science.aat1168

Goldford, Joshua E.; Lu, Nanxi; Bajić, Djordje; Estrela, Sylvie; Tikhonov, Mikhail; Sanchez-Gorostiaga, Alicia; Segrè, Daniel; Mehta, Pankaj; Sanchez, Alvaro (August 2018, Science)

Full Text Available
Remnants of an Ancient Metabolism without Phosphate

https://doi.org/10.1016/j.cell.2017.02.001

Goldford, Joshua E.; Hartman, Hyman; Smith, Temple F.; Segrè, Daniel (March 2017, Cell)

Full Text Available
Genome-Scale Architecture of Small Molecule Regulatory Networks and the Fundamental Trade-Off between Regulation and Enzymatic Activity

https://doi.org/10.1016/j.celrep.2017.08.066

Reznik, Ed; Christodoulou, Dimitris; Goldford, Joshua E.; Briars, Emma; Sauer, Uwe; Segrè, Daniel; Noor, Elad (September 2017, Cell Reports)

Full Text Available
A roadmap for the functional annotation of protein families: a community perspective

https://doi.org/10.1093/database/baac062

de Crécy-lagard, Valérie; Amorin de Hegedus, Rocio; Arighi, Cecilia; Babor, Jill; Bateman, Alex; Blaby, Ian; Blaby-Haas, Crysten; Bridge, Alan J.; Burley, Stephen K.; Cleveland, Stacey; et al (August 2022, Database)

Abstract Over the last 25 years, biology has entered the genomic era and is becoming a science of ‘big data’. Most interpretations of genomic analyses rely on accurate functional annotations of the proteins encoded by more than 500 000 genomes sequenced to date. By different estimates, only half the predicted sequenced proteins carry an accurate functional annotation, and this percentage varies drastically between different organismal lineages. Such a large gap in knowledge hampers all aspects of biological enterprise and, thereby, is standing in the way of genomic biology reaching its full potential. A brainstorming meeting to address this issue funded by the National Science Foundation was held during 3–4 February 2022. Bringing together data scientists, biocurators, computational biologists and experimentalists within the same venue allowed for a comprehensive assessment of the current state of functional annotations of protein families. Further, major issues that were obstructing the field were identified and discussed, which ultimately allowed for the proposal of solutions on how to move forward.
more » « less

Search for: All records